Papers with evaluation paradigms
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement (P18-1)
Copied to clipboard
| Challenge: | a number of post hoc explanation methods for deep neural networks have been proposed . due to the complexity of the DNNs they explain, these methods are necessarily approximations and come with their own sources of error. |
| Approach: | They propose two evaluation paradigms that cover two important classes of NLP problems . they propose LIMSSE, LRP and DeepLIFT as the most effective explanation methods . |
| Outcome: | The proposed methods are most effective for explaining deep neural networks in NLP . the proposed methods can explain complex models without manual annotation . |
Are Neural Topic Models Broken? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm . |
| Approach: | They show that neural topic models fare worse in both respects compared to an established classical method. |
| Outcome: | The proposed method outperforms the members of the ensemble in both respects. |
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests (2025.findings-emnlp)
Copied to clipboard
Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi
| Challenge: | MMLU and BBH are three evaluation paradigms for language learning models . interactive games are superior to standard benchmarks in discriminating models based on human cognitive assessments . |
| Approach: | They examine three evaluation paradigms: standard benchmarks, interactive games and cognitive tests . they examine whether interactive games are more effective at discriminating LLMs . |
| Outcome: | The results show that interactive games are superior to standard benchmarks in discriminating models. |